Papers with Multimodal Large Language Models
Copied to clipboard
| Challenge: | Recent studies improve visual contrastive decoding (VCD) by constructing more informative auxiliary views. |
| Approach: | They propose to construct an object-aligned auxiliary view that disrupts unsupported tokens and produces a stronger contrast signal. |
| Outcome: | Empirically, the proposed method shows consistent gains on two popular object hallucination benchmarks across two MLLMs. |
Copied to clipboard
| Challenge: | Existing AVR benchmarks focus on single-step reasoning, emphasizing the end result but neglecting the multi-stage nature of reasoning process. |
| Approach: | They propose a multi-stage AVR benchmark based on RAVEN to assess reasoning across varying levels of complexity. |
| Outcome: | The proposed metric considers the correctness of intermediate steps in addition to the final outcomes. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) struggle with identifying and categorizing student errors in multimodal mathematical contexts. |
| Approach: | They propose a new framework that decomposes error detection into three phases with specialized agents. |
| Outcome: | The proposed framework shows higher accuracy in error step identification and 3% improvement in error categorization on real-world educational data. |
Copied to clipboard
| Challenge: | Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development. |
| Approach: | They introduce a region-based score to quantify a dataset's reliance on global versus local visual information. |
| Outcome: | The proposed model-based score systematically compares model performance on image patches versus full images to determine if tasks require holistic image understanding or can be solved with partial or localized visual cues. |
Copied to clipboard
| Challenge: | SceMQA focuses on core science subjects including Mathematics, Physics, Chemistry, and Biology. |
| Approach: | They propose to use SceMQA to evaluate multimodal question answering at college entrance level. |
| Outcome: | The proposed model provides specific knowledge points for each problem and detailed explanations for each answer. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for Multimodal Large Language Models (MLLMs) are inadequate to assess their robustness to irrelevant or distracting visual context. |
| Approach: | They propose a patch-context-robustness index to measure MLLMs' robustness to visual context variations. |
| Outcome: | The proposed score measures the robustness of MLLMs to visual contexts across 15 vision-language benchmarks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have demonstrated proficiency in handling a variety of visual-language tasks, but their ability to extrapolate from image sequences has been less investigated. |
| Approach: | They propose a new benchmark to assess MLLMs’ sequential image reasoning abilities. |
| Outcome: | The proposed benchmark features 4,761 diverse image sequences with varying lengths. |
Copied to clipboard
| Challenge: | Existing music-focused benchmarks are fragmented, largely single-modality, Western-centric . existing methods for evaluating MLLMs are lacking reproducibility and reliability . |
| Approach: | They propose to develop a musically multimodal benchmark that will integrate music into the benchmark. |
| Outcome: | The proposed benchmark will integrate culturally diverse musical material beyond the dominant Western canon. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are used for document information extraction, but their impact on document information processing remains unclear. |
| Approach: | They propose an automated hierarchical error analysis framework that leverages large language models to diagnose errors systematically. |
| Outcome: | The proposed framework can achieve comparable performance to OCR-enhanced approaches. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models struggle with Long Video Understanding due to their limited context window and the distributed nature of salient information across many redundant frames. |
| Approach: | They propose a training framework that mimics a human reasoning process to train Long Video Understanding models. |
| Outcome: | The proposed framework achieves 77.6% performance on Video MME, LongVideo, and MLVU benchmarks while yielding 5% improvement on Llama 4 Scout. |
Copied to clipboard
| Challenge: | Existing text-to-image models excel at generating high-quality object-centric images from instructions, but lack of data for complex interactions. |
| Approach: | They propose a multimodal Large Language Models-generated dataset to benchmark and enhance interaction-rich images. |
| Outcome: | The proposed approach improves image quality and automatic and human evaluations show improvements. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have demonstrated proficiency in diverse tasks across different domains. |
| Approach: | They propose a method that integrates multimodal instruction tuning with Conditional Mixture-of-LoRA. |
| Outcome: | Experimental results show that MixLoRA outperforms LoRA with the same or higher ranks . MLLMs have demonstrated remarkable proficiency in diverse tasks across domains . |
Copied to clipboard
| Challenge: | Recent advances in Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal reasoning abilities, but also raised new safety concerns, especially when faced with adversarial multimodal inputs. |
| Approach: | They propose a modular and adaptive inference-time intervention technology, AutoSteer, that integrates a safety awareness score, an adaptive safety prober, and a lightweight Refusal Head to modulate generation when safety risks are detected. |
| Outcome: | Experiments on LLaVA-OV and Chameleon show that AutoSteer significantly reduces the Attack Success Rate (ASR) for textual, visual, and cross-modal threats while maintaining general abilities. |
Copied to clipboard
| Challenge: | Unlike professional Business-to-Consumer (B2C) e-commerce platforms, consumer-to consumer (C2C), is mainly targeting individual sellers. |
| Approach: | They develop an intelligent product listing tool that generates product descriptions using various product attributes such as category, brand, color, condition, etc. |
| Outcome: | The proposed tool outperforms the base model in domain-specific tasks while producing less hallucination. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) integrate visual and textual inputs, yet modality alignment remains one of the most challenging aspects. |
| Approach: | They propose a token-level supervision alignment method that enables more precise visual-text alignment during pretraining. |
| Outcome: | The proposed method improves performance across various model sizes, with smaller models benefiting the most. |
Copied to clipboard
| Challenge: | a new multimodal decision-making benchmark evaluates the integrated capabilities of multimodal large language models. |
| Approach: | They propose a multimodal decision-making benchmark for evaluating MLLMs . they propose an automatic evaluation protocol to assess 10 prevalent ML models . |
| Outcome: | The proposed benchmark improves performance of multimodal large language models in three scenarios . the model is required to integrate multiple capabilities to make accurate decisions . |
Copied to clipboard
| Challenge: | Music audio-visual question answering presents unique challenges with dense audio-visual content, intricate temporal dynamics, and the need for domain-specific knowledge. |
| Approach: | They analyze Music AVQA datasets and analyze their results to identify key design patterns . they propose concrete future directions for incorporating musical priors . |
| Outcome: | The proposed architectures are critical for success in Music AVQA, the authors argue . they suggest concrete future directions for incorporating musical priors . |
Copied to clipboard
| Challenge: | Existing models merging methods often lead to suboptimal performance due to harmful models . et al., 2018; 59: 59-64. |
| Approach: | They propose an uncertainty-guided MLLM merging algorithm that integrates models into a single MLML. |
| Outcome: | The proposed algorithm improves on held-in and held-out vision-language benchmarks. |
Copied to clipboard
| Challenge: | Existing approaches to VideoQA often fail when complex reasoning or temporal relationships are involved. |
| Approach: | They propose a method that leverages reasoning processes generated by Multimodal Large Language Models to improve VideoQA models. |
| Outcome: | The proposed method improves VideoQA models on three benchmarks. |
Copied to clipboard
| Challenge: | Current AM methods focus on extracting attributes from unimodal text, underutilizing multimodal data. |
| Approach: | They propose a framework for multimodal self-correction instruction tuning to extract new attributes from images and text with Multimodal Large Language Models. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on two datasets. |
Copied to clipboard
| Challenge: | Existing video moment retrieval methods rely on sparse frame sampling, risking information loss. |
| Approach: | a new video-based framework enhances memory efficiency while maintaining high information resolution . SMORE uses query-guided captions to encode semantics aligned with user intent . |
| Outcome: | a new framework improves memory efficiency while maintaining high information resolution . it achieves state-of-the-art performance on QVHighlights, Charades-STA, and ActivityNet-Captions benchmarks . |
Copied to clipboard
| Challenge: | Multimodal Large Language Models are pre-trained on image-text caption data and interleaved document data. |
| Approach: | They propose to train an efficient MLLM as a Unified Mulitmodal Data Quality Classifier to filter image-text caption and interleaved data. |
| Outcome: | The proposed method enables efficient creation of sample-score pairs for caption and interleaved data to train UniFilter. |
Copied to clipboard
| Challenge: | Existing methods for question decomposition focus on unimodal language models, but question decomposing capability of Multimodal Large Language Models (MLLMs) has yet to be explored. |
| Approach: | They propose a finetuning dataset and a training objective for selective decomposition to enhance the model's question decomposing capability. |
| Outcome: | The proposed dataset shows that existing models struggle to produce high-quality sub-questions. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have seen remarkable progress in providing general instruction-following ability, but struggle with critical problems when required to provide a detailed and accurate response to a visual instruction. |
| Approach: | They propose to enhance the mapping process by using retrieval-augmented tag tokens, which contain rich object-aware information such as object names and attributes. |
| Outcome: | The proposed model outperforms baselines that share the same language model and training data on 12 benchmarks and shows zero-shot capability when provided with specific datastores. |
Copied to clipboard
| Challenge: | Existing MLLMs are optimized for single-task scenarios and struggle to generalize to diverse contexts. |
| Approach: | They propose a framework that integrates multitask reinforcement learning and generalization capabilities of MLLMs to optimize the judge model across multiple tasks. |
| Outcome: | The proposed framework outperforms baseline models in judgment consistency and correlation with human preferences. |
Copied to clipboard
| Challenge: | Multimodal large language models are increasingly used for movie understanding . however, their performance on movies lags behind other video understanding tasks . |
| Approach: | They analyze movie knowledge, cinematographic knowledge, and critical analysis to identify where MLLMs fail . ML models are increasingly used for movie understanding . |
| Outcome: | The results show that MLLMs outperform existing methods in small-scale settings involving factual knowledge but fail when cinematographic and critical analysis is required. |
Copied to clipboard
| Challenge: | Existing approaches to MCIT address Catastrophic Forgetting and Knowledge Transfer (KT) but using a fixed number of shared LoRA blocks across tasks can lead to knowledge interference. |
| Approach: | They propose a framework that uses a fixed number of shared LoRA blocks to reduce knowledge interference. |
| Outcome: | The proposed framework outperforms existing approaches on the latest MCIT benchmark. |
Copied to clipboard
| Challenge: | Video-guided machine translation (VMT) aims to improve translation quality by integrating contextual information from paired short video clips. |
| Approach: | They propose a plug-and-play framework for video-guided machine translation with multimodal large language models. |
| Outcome: | The proposed framework improves performance of MLLMs while reducing computational cost. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models have shown significant promise in various applications, but a comprehensive evaluation of their long-context capabilities remains underexplored. |
| Approach: | They propose a benchmark to assess the long-context capabilities of multimodal large language models. |
| Outcome: | The proposed benchmark compared MLLMs with API-based and open-source models in a long-context scenario. |
Copied to clipboard
| Challenge: | Recent approaches demonstrate that MLLMs can be adapted into competitive embedding models via large-scale contrastive learning. |
| Approach: | They propose a compressed pre-training phase which serves as a warm-up stage for contrastive learning. |
| Outcome: | The proposed model achieves state-of-the-art among MLLMs of comparable size on the MMEB, realizing optimization in both efficiency and effectiveness. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models lack general structure understanding abilities for text-rich document images. |
| Approach: | They propose to use unified structure learning to boost the performance of MLLMs by encoding structure information into text-rich images. |
| Outcome: | The proposed model achieves state-of-the-art on 10 visual document understanding benchmarks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have been gaining popularity in multimodal tasks . a bilingual benchmark is available for MLLM users to evaluate their multimodal capabilities . |
| Approach: | They propose a bilingual multimodal ability norms benchmark that measures multimodality across nine tasks. |
| Outcome: | The proposed benchmark compared human performance against state-of-the-art MLLMs. |
Copied to clipboard
| Challenge: | despite significant strides in multimodal tasks, MLLMs are plagued by the critical issue of hallucination. |
| Approach: | They propose a meta-evaluation benchmark to facilitate evaluation of advancements in hallucination detection methods. |
| Outcome: | The proposed framework validates hallucinations robustly and provides strategic insights . MHaluBench is a meta-evaluation benchmark designed to facilitate evaluation . |
Copied to clipboard
| Challenge: | Existing studies show that multimodal large language models can learn from text-image data. |
| Approach: | They propose to train multimodal large language models on large amounts of text-image data . they also show a boost in few-shot learning performance across various multilingual tasks . |
| Outcome: | The proposed dataset is not public and is only in English . it is the first large-scale multilingual and multimodal document corpus crawled from the web. |
Copied to clipboard
| Challenge: | Existing evaluation methods for mobile GUI agents rely on static frame assessments or offline static apps. |
| Approach: | They propose an evaluation system that leverages large language models as reward models to verify task completion and process achievement. |
| Outcome: | The proposed system addresses the limitations of traditional function based evaluation methods on online dynamic apps. |
Copied to clipboard
| Challenge: | Existing document understanding models focus on key information and generate answers straightforwardly . existing models ignore evidence from source documents and lack interpretability . |
| Approach: | They propose a visual encoder that fuses text into visual encoded visual encodes . they use multimodal large language models as data generators and checkers to generate step-wise question-and-answer pairs for document images. |
| Outcome: | The proposed model can answer step-wise questions without compromising the performance of the original model. |
Copied to clipboard
| Challenge: | Recent efforts to accelerate inference in Multimodal Large Language Models have focused on visual token compression. |
| Approach: | They propose a framework that leverages downsampling as a discriminator to denoise existing benchmarks. |
| Outcome: | The proposed evaluation framework leverages downsampling as a discriminator to denoise existing benchmarks. |
Copied to clipboard
| Challenge: | Existing methods for person anomaly search fail to address the complexities of real-world security, authors say . Existing approaches fail to detect subtle semantic distinctions, authors argue . |
| Approach: | They propose a framework that decouples retrieval into two stages . structure-aware coarse retrieval and detective squad interaction are proposed . |
| Outcome: | The proposed framework achieves state-of-the-art performance by balancing efficiency and semantic reasoning. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) often rely on spurious correlations, undermining their robustness and generalization. |
| Approach: | They propose a causal mediation-based debiasing framework to address correlation bias in MLLMs . they distinguish core semantics from spurious textual and visual contexts using counterfactual examples . |
| Outcome: | The proposed framework surpasses existing state-of-the-art models on sarcasm detection and sentiment analysis tasks. |
Copied to clipboard
| Challenge: | Existing RAG research focuses on textual data, overlooking rich visual content in financial documents. |
| Approach: | They propose a visual RAG benchmark tailored for finance that integrates multimodal data and provides visual citation to ensure traceability. |
| Outcome: | The proposed visual RAG benchmark integrates multimodal data and provides visual citation to ensure traceability. |
Copied to clipboard
| Challenge: | Existing approaches to search for images using single-modality are limited by representation space fragmentation. |
| Approach: | They propose a unified representation framework that achieves efficient query-target alignment . they introduce a multi-level Chain-of-Thought prompting strategy that guides MLMs to generate discriminative, semantically compatible captions for target images . |
| Outcome: | The proposed framework achieves efficient query-target alignment through synergistic components. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) exhibit remarkable performance across a wide range of domains. |
| Approach: | They propose a multimodal prompt tuning approach for efficient instruction tuning of MLLMs. |
| Outcome: | The proposed approach shows superior performance on multimodal evaluation datasets compared to state-of-the-art methods. |
Copied to clipboard
| Challenge: | Existing methods for hallucination mitigation rely on external verification or post-hoc correction, lacking internal mechanism to validate outputs directly during training. |
| Approach: | They propose a unified closed-loop training framework that encourages multimodal consistency for cross-modal understanding in MLLMs. |
| Outcome: | The proposed framework encourages multimodal consistency for cross-modal understanding in MLLMs. |
Copied to clipboard
| Challenge: | Existing multimodal large language models suffer from systematic failures in basic visual understanding. |
| Approach: | They propose a tool-augmented reasoning framework with three targeted compensation strategies to address these problems. |
| Outcome: | The proposed framework improves visual grounding by re-injecting the original image to mitigate visual forgetting, the authors show . the proposed framework also improves the accuracy of the visual inputs, the researchers show - and the results are promising . |
Copied to clipboard
| Challenge: | Existing methods to extrapolate and comprehend changes in object states are limited . relying on a small set of symbolic words to represent changes has restricted expressiveness of language. |
| Approach: | They propose a dataset and benchmark to evaluate multimodal large language models . they investigate causal relations between a concrete action and the change . |
| Outcome: | The proposed method achieves near parity with GPT-4V ratings across helpfulness, accuracy, reasoning, and other key metrics. |
Copied to clipboard
| Challenge: | Tabular data is often captured in image form across a wide range of real-world scenarios. |
| Approach: | They propose a framework that enables MLLMs to answer queries over large tables. |
| Outcome: | The proposed framework outperforms existing methods by 7.0% in retrieval recall and 6.1% in answer accuracy on a newly constructed dataset with 48,504 unique tables. |
Copied to clipboard
| Challenge: | Existing backdoor attacks on Multimodal Large Language Models are less applicable to open-ended conversations with users. |
| Approach: | They propose a shadow-activated backdoor attack scenario where attackers inject malicious content into the responses of MLLMs when the responses explicitly relate to the shadowed object. |
| Outcome: | The proposed framework achieves the desired behaviors by constructing a poisoned dataset and implementing an attention-regularized tuning strategy. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models suffer from hallucinations, especially errors in object existence, attributes, or relations. |
| Approach: | They propose a framework that decomposes responses into atomic queries and estimates confidence using self-consistency or self-confidence aggregation. |
| Outcome: | Experiments on five benchmarks show that TACO outperforms direct prompting and Visual Contrastive Decoding and improves confidence calibration. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) can identify the orientation of input images rotated 0°, 90°, 180°, and 270°. |
| Approach: | They propose a manually-filtered benchmark to evaluate MLLMs' ability to accurately identify rotation in input images. |
| Outcome: | The proposed model improves on the 'rotational cues' of 360° and 180° images, but not 90° and 270° rotations. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) show impressive capabilities across visual–language tasks, but their capacity to evaluate artistic expression remains limited. |
| Approach: | They propose an attribute-specific multi-LoRA approach where each attribute corresponds to a distinct evaluation dimension in the scoring rubric. |
| Outcome: | The proposed approach increases correlation from 0.468 to 0.653 on Qwen2.5-VL-7B, with the largest gains on perceptual dimensions and narrowed gaps on higher-order attributes. |
Copied to clipboard
| Challenge: | Speech signals convey abundant speaker-related metadata, yet current privacy research focuses on identity-centric voiceprint protection, leaving sensitive Speaker Attribute Privacy (SAP) underexplored. |
| Approach: | They propose a large-scale Chinese dataset to evaluate speaker-related privacy leakage . the dataset includes 227.3 hours of audio from 1,000 speakers . |
| Outcome: | The proposed model systematically evaluates speaker-related privacy leakage in everyday scenarios. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are limited by context length when processing long videos. |
| Approach: | They propose a training-free method that flexibly reduces redundancy by allocating compression ratios among time and model layers with theoretical guarantees. |
| Outcome: | Experiments on videoMME, MLVU, LongVideoBench, and LVBench show that AdaRETAKE outperforms existing methods by 2.3% and 2.8% for 7B and 72B models. |
Copied to clipboard
| Challenge: | Existing multimodal large language models lack the ability to perceive the visual world with a deep concept structure cognition. |
| Approach: | They propose a concept-level benchmark to assess MLLMs’ hierarchical concept understanding and reasoning abilities. |
| Outcome: | The proposed model outperforms state-of-the-art models in concept structure reasoning evaluation. |
Copied to clipboard
| Challenge: | Existing MLLMs lack robustness in multimodal causal reasoning compared to their performance in textual settings. |
| Approach: | They propose a novel multimodal chain-of-thought (CoT) reasoning benchmark that leverages siamese images and text pairs to challenge MLLMs. |
| Outcome: | The proposed benchmark leverages siamese images and text pairs to challenge MLLMs. |
Copied to clipboard
| Challenge: | Prior work focused on typographic and pixel-level perturbations, leaving the study of SCO unexplored. |
| Approach: | They propose a framework that exploits MLLMs' diagrammatic reasoning capabilities to bypass safety guardrails. |
| Outcome: | The proposed framework exploits the model's reasoning capabilities to bypass safety guardrails. |
Copied to clipboard
| Challenge: | Large Language Models and Multimodal Large Language Modells can memorize sensitive information, raising ethical and privacy concerns. |
| Approach: | They propose a novel unlearning framework that selectively clips neurons based on their relative importance to the targeted forget data. |
| Outcome: | The proposed framework selectively clips neurons based on their relative importance to the targeted forget data, curated for different modalities. |
Copied to clipboard
| Challenge: | Current benchmarks focus on coarse-grained knowledge, leaving the intricacies of fine-grounded knowledge unexplored. |
| Approach: | They propose a benchmark and dataset specifically designed for FG multimodal entity knowledge editing. |
| Outcome: | The proposed benchmark underscoring the complexity of FG knowledge editing in MLLMs. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) provide no visibility into which parts of visual data informed their conclusions. |
| Approach: | They propose a semi-automatic approach to attribute reasoning process by highlighting regions in charts and graphs that justify model answers. |
| Outcome: | The proposed method improves attribution accuracy by up to 15 percentage points compared to baseline methods and achieves high semantic similarity with ground truth responses. |
Copied to clipboard
| Challenge: | Existing work focuses on generating citations for text-only content . experimental results reveal MLLMs struggle to ground outputs reliably when handling multimodal input . |
| Approach: | They propose a benchmark to assess the ability of MLLMs to generate text with citations in multimodal contexts. |
| Outcome: | The proposed benchmark assesses the ability of MLLMs to generate text with citations in multimodal contexts. |
Copied to clipboard
| Challenge: | Existing benchmarks and evaluation protocols suffer from inadequate or homogeneous creation of key points, exorbitant cost of data creation, and limited evaluation scopes. |
| Approach: | They propose an automatic framework which leverages Monte Carlo Tree Search to construct numerous and diverse descriptive sentences that thoroughly represent video content in an iterative way. |
| Outcome: | The proposed framework improves MCTS-VCB and DREAM-1K on video captioning tasks by 25.0% and 16.3% respectively. |
Copied to clipboard
| Challenge: | Automated Essay Scoring (AES) systems face three major challenges: reliance on handcrafted features that limit generalizability, difficulty in capturing fine-grained traits like coherence and argumentation, and inability to handle multimodal contexts. |
| Approach: | They propose a multimodal benchmark to evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
| Outcome: | The proposed system can evaluate AES capabilities across lexical-, sentence-, and discourse-level traits without manual feature engineering. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have shown impressive capabilities in vision-language understanding but their visual input remains fixed throughout the reasoning process. |
| Approach: | They propose a model-agnostic tree search algorithm tailored for vision-level reasoning that allows MLLMs to explore textual tokens while visual input remains fixed throughout reasoning process. |
| Outcome: | The proposed algorithm outperforms strong large models such as GPT-4o on high-resolution benchmarks and improves performance on a series of elaborate high-level benchmarks. |
Copied to clipboard
| Challenge: | Existing efforts to improve task accuracy or enrich COT generation are lacking in multimodal large language models. |
| Approach: | They propose a Faithful-First Reasoning, Planning, and Acting framework that evaluates faithfulness of intermediate reasoning and uses it to plan and execute faithfulness-aware actions during inference. |
| Outcome: | The proposed framework improves perceptual faithfulness by up to 24% over prompt-based and tool-augmented reasoning frameworks without degrading task accuracy. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) show reasoning promise, yet their visual perception is a bottleneck. |
| Approach: | They propose a visual perception benchmark to test the visual perception of MLLMs. |
| Outcome: | The proposed benchmark examines MLLMs' visual perception abilities with 1758 images and 2612 questions. |
Copied to clipboard
| Challenge: | Existing pruning methods fail to account for unique token attributes across layers and modalities inherent to MLLMs. |
| Approach: | They propose a pruning framework that takes into account unique token attributes across layers and modalities inherent to MLLMs. |
| Outcome: | The proposed pruning framework outperforms existing pruning techniques on two state-of-the-art MLLMs. |
Copied to clipboard
| Challenge: | Existing egocentric video datasets do not support the personalization and long-context reasoning required for episodic memory retrieval. |
| Approach: | They propose a benchmark framework that uses MLLMs and reflective Chain-of-Thought to ground user queries in personal memory explicitly. |
| Outcome: | The proposed framework outperforms state-of-the-art benchmarks on three benchmarks . it can be used to generate detailed target video descriptions in long-context contexts based on user-specific object annotations enriched with user-specified object annotation data . |
Copied to clipboard
| Challenge: | Existing open-source MLLMs fail to fully capture dense information embedded in charts . current models still face significant challenges in understanding and analyzing visual tasks such as captioning and question answering. |
| Approach: | They propose a chart-to-code MLLM which leverages Code LLMs as the language backbone to enhance the executability of the generated code. |
| Outcome: | The proposed model surpasses existing open-source models on chart-to-code benchmarks with only 7B parameters and provides lossless representations that contain all critical details. |
Copied to clipboard
| Challenge: | Existing attacks focus on increasing the complexity of the modified visual task and do not explicitly leverage the model’s own reasoning incentives. |
| Approach: | They propose a framework that decomposes and reassembles harmful visual semantics and constructs a gamified scene that drives the model to explore, reconstruct intent and answer as part of winning the game. |
| Outcome: | Experiments on reasoning and non-reasoning MLLMs show that the proposed framework outperforms baseline models on both vision and text. |
Copied to clipboard
| Challenge: | Multipanel images are a common form of visual representations, and humans can achieve approximately 99% accuracy on these questions. |
| Approach: | They propose a benchmark that tests multipanel visual reasoning models with 6,600 triplets of questions, answers, and multipanel images. |
| Outcome: | The proposed benchmark features 6,600 triplets of questions, answers, and multipanel images that challenge state-of-the-art Multimodal Large Language Models (MLLMs) human users can attain approximately 99% accuracy on these questions, compared with previous benchmarks. |
Copied to clipboard
| Challenge: | Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces. |
| Approach: | They propose a benchmark to evaluate MLLMs' understanding of multimodal social media content and a large-scale YouTube tagging dataset to evaluate their performance. |
| Outcome: | The proposed model performs better in a zero-shot setting, suggesting potential improvements. |
Copied to clipboard
| Challenge: | Recent advances in machine learning (MU) have enabled the selective removal of private or sensitive information encoded within deep neural networks. |
| Approach: | They propose to "reformulate" the task of multimodal MU in the era of MLLMs by preserving only the visual patterns associated with a given entity while preserving the corresponding textual knowledge. |
| Outcome: | The proposed method surpasses baselines that finetuned MLLMs with VQA data directly through Gradient Ascent (GA) or Negative Preference Optimization (NPO), across all evaluation dimensions. |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating MLLMs have not addressed active perception . a novel benchmark is proposed to evaluate active perception in ML models . |
| Approach: | They propose a benchmark to evaluate active perception in Multimodal Large Language Models . they restrict the perceptual field of a model and require it to actively zoom or shift it . |
| Outcome: | The proposed benchmark focuses on a specialized form of Visual Question Answering (VQA) that eases and quantifies the evaluation yet challenging for existing MLLMs. |
Copied to clipboard
| Challenge: | Existing models struggle to maintain stable understanding performance and low GPU memory overhead. |
| Approach: | They propose a training-free architecture for real-time and accurate understanding of video streams . HERMES reuses a compact KV cache, enabling efficient streaming understanding . |
| Outcome: | The proposed architecture achieves 10 faster TTFT compared to prior SOTA. |
Copied to clipboard
| Challenge: | Existing multimodal large language models incorporate visual and textual information, but introduces new and complex safety risks. |
| Approach: | They propose a safety reasoning framework that integrates visual modalities into multimodal models to help them resist jailbreak attacks. |
| Outcome: | The proposed framework improves model safety while avoiding over-defense . it is based on a large-scale safety reasoning dataset . |
Copied to clipboard
| Challenge: | Multimodal Large Language Models are increasingly used in Personalized Image Aesthetic Assessment (PIAA) however, their predictions may reflect subtle biases influenced by demographic factors such as gender, age, and education. |
| Approach: | They propose to evaluate MLLMs along two complementary dimensions: (1) stereotype bias and (2) alignment between model outputs and genuine human aesthetic preferences. |
| Outcome: | The proposed benchmark covers three subtasks: aesthetic perception, assessment, empathy and alignment between outputs and genuine human aesthetic preferences. |
Copied to clipboard
| Challenge: | Existing MLLMs have a visual question answering capability but lack domain-specific information. |
| Approach: | They propose a framework for language model modules in MLLMs when handling projected image features and verify this hypothesis using logit lens. |
| Outcome: | The proposed framework will yield a 10% change in accuracy at most, shedding light on the development of cross-domain, all-encompassing MLLMs in the future. |
Copied to clipboard
| Challenge: | Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios. |
| Approach: | They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity. |
| Outcome: | The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues. |
Copied to clipboard
| Challenge: | Recent studies have introduced eclectic strategies to enhance MLLMs’ reasoning capabilities, but they remain related to a single language. |
| Approach: | They propose a modular approach that instructs models to abstract key elements of the reasoning process and refine reasoning trajectories via self-correction. |
| Outcome: | The proposed approach improves multimodal reasoning, gets aligned performances among the languages approaching strong models and improves the model's performance. |
Copied to clipboard
| Challenge: | Text-Centric Visual Question Answering (TEC-VQA) is a text-centric visual task understanding tool. |
| Approach: | They introduce a benchmark that features human expert annotations across 9 languages . they prioritize the text in question-answer pairs while disregarding visual text in images . |
| Outcome: | The proposed benchmarks prioritize the text in question-answer pairs while disregarding visual text in images. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) often hallucinate due to fragile, linear reasoning and weak visual grounding. |
| Approach: | They propose a framework that reformulates reasoning as a hierarchical search with self-verification and replaces linear Chain-of-Thought with a tree-search policy capable of backtracking to correct logical errors. |
| Outcome: | The proposed framework outperforms state-of-the-art methods on hallucination and safety benchmarks. |
Copied to clipboard
| Challenge: | Existing research indicates that even state-of-the-art MLLMs still suffer from some straightforward visual question-answering (VQA) problems. |
| Approach: | They propose to use a model-based benchmark to investigate model laziness to identify models that err when answering simple visual questions about an image. |
| Outcome: | The proposed model laziness is found to be widespread in current MLLMs, including GPT-4o, Gemini-1.5-pro, Claude 3, LLaVA-1.5, LLva-1.6, and QWen-VL. |
Copied to clipboard
| Challenge: | Existing knowledge editing paradigms suffer from editing decoupling failures . entity knowledge is sequestered into disentangled modality-specific pathways . |
| Approach: | They propose a method that explicitly disentangles and localizes modality-specific neuron groups for targeted knowledge. |
| Outcome: | The proposed method outperforms baselines in reliability and consistency while preserving model locality. |
Copied to clipboard
| Challenge: | Existing benchmarks for evaluating instruction-following capabilities focus on verbal instructions in the textual modality. |
| Approach: | They propose to incorporate vision-dependent constraints into instruction design to enable a more rigorous assessment of how well MLLMs align their outputs with both visual input and textual instructions. |
| Outcome: | The proposed benchmark incorporates vision-dependent constraints into instruction design, enabling a more rigorous and fine-grained assessment of how well MLLMs align their outputs with both visual input and textual instructions. |
Copied to clipboard
| Challenge: | Existing MLLM benchmarks and unified evaluation frameworks cannot accurately and efficiently reflect the ability of MLMLs. |
| Approach: | They propose a semi-automated benchmark curated using a pipeline that filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding. |
| Outcome: | The proposed benchmark reduces the number of samples by 76% and evaluation time by 77% while it can more effectively distinguish different models’ abilities. |
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models have remained opaque. |
| Approach: | They propose a method to convert dense MLLMs into fine-grained Mixture-of-Experts architectures. |
| Outcome: | The proposed method outperforms random expert pruning and sparse activation and model pruning. |
Copied to clipboard
| Challenge: | Existing parameter-efficient approaches to multimodal Continual Instruction Tuning suffer from knowledge interference and inefficient capacity expansion, limiting scalability. |
| Approach: | They propose a framework for multimodal Continual instruction tuning that decomposes adaptation weights into a globally shared pool of orthonormal bases to capture task-invariant knowledge. |
| Outcome: | Experiments show that MoBLoRA outperforms state-of-the-art methods while maintaining superior parameter efficiency. |
Copied to clipboard
| Challenge: | Recent research shows that multimodal large language models are vulnerable to jailbreak attacks . |
| Approach: | They propose a jailbreak attack method based on auto-generated flowcharts . the flowchartings are then combined with a benign textual prompt to execute the attack . |
| Outcome: | The proposed method achieves an attack success rate of up to 96% via images and 78% via videos across multiple MLLMs. |
Copied to clipboard
| Challenge: | Recent advances in multimodal large language models (MLLMs) have demonstrated exceptional capabilities in visual perception and understanding, but they also suffer from hallucinations, which limit their reliability as AI systems. |
| Approach: | They propose a benchmark to evaluate self-awareness in perception for multimodal large language models (MLLMs) by integrating image information with knowledge quadrants, and propose MM-SAP to evaluate this capability. |
| Outcome: | The proposed benchmark offers detailed analysis of MLLMs with self-awareness in perception. |
Copied to clipboard
| Challenge: | Recent advances in Multimodal Large Language Models have led to a significant surge in the resource consumption of these models. |
| Approach: | They propose a method to reduce image tokens using visual query data by using CLIP metrics to reduce computational overhead and maintain consistent performance. |
| Outcome: | The proposed method has been extensively tested across 12 datasets and shows a significant reduction in computational overhead while maintaining a consistent level of performance. |
Copied to clipboard
| Challenge: | Existing approaches to generate SVG-based fonts struggle with semantic ambiguity and inefficiency . edward mcginley: generic text tokenizers fragment coordinate-dense SVG XML into excessively long sequences . |
| Approach: | They propose a system that treats SVG generation as a conditional language modeling task . they propose linguistic supervision framework that decomposes typographic style into interpretable linguistic dimensions . |
| Outcome: | The proposed system improves CLIP score by +23% while reducing geometric error by 48% and boosts generation efficiency by 18% Command-per-Token (C/T) ratio. |
Copied to clipboard
| Challenge: | Existing MLLMs still struggle to achieve precise grounding in multi-image scenarios. |
| Approach: | They propose a Chain-of-Thought framework that integrates single-image grounding with multi-image comprehension to address this challenge. |
| Outcome: | The proposed model outperforms existing models in multi-image grounding tasks by 24.94% and surpasses larger 70B models. |
Copied to clipboard
| Challenge: | Existing jailbreak methods only use a single image, restricting the attack space . Existing frameworks only use single image to distribute harmful requests across multiple images . |
| Approach: | They propose a compositional jailbreak framework that leverages Distributed instruction, Multimodal evidence and a Number chain task to fully enhance the jailbreak performance. |
| Outcome: | The proposed framework achieves attack success rates of over 90% on GPT-4o, Gemini-2.5-pro and Claude Sonnet 4 . |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have advanced visual and language understanding, but their potential in Chinese Classical Studies (CCS) remains underexplored due to the lack of specialized benchmarks. |
| Approach: | They propose to develop a multimodal benchmark specifically designed for Chinese Classical Studies across multiple subdomains to bridge this gap. |
| Outcome: | The proposed benchmark spans seven core subdomains with a total of 45 meticulously designed tasks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have potential for cross-modal understanding . but extending MLLM to handle diverse modalities introduces two challenges . |
| Approach: | They propose a dual-stage compression mechanism to reduce the number of modality tokens per modality and condense it into a single, compact token sequence. |
| Outcome: | Experiments show that Flex-M3 outperforms its counterpart trained on only full-modality data. |
Copied to clipboard
| Challenge: | MLLMs lack visual grounding mechanism to read text embedded in images, or rely on parametric shortcuts . despite strong OCR capabilities, models suffer performance degradation of 12.7% in the VQ setting . |
| Approach: | They propose a plug-and-play training strategy that invalidates shortcuts in text prompts . they propose 'vq' setting where text queries are rendered directly onto images . |
| Outcome: | The proposed training strategy surpasses the base model by 5.4% and GRPO based on original images by 2.7% on four representative OOD benchmarks. |
Copied to clipboard
| Challenge: | Existing models for GUI understanding ignore a key GUI-referring task: screen reading based on user-indicated points. |
| Approach: | They propose a Tree-of-Lens agent that constructs a Hierarchical Layout Tree based on user input points and a GUI screenshot. |
| Outcome: | The proposed agent can interpret the Screen Point-and-Read task on mobile, web, and operating systems. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) trained on massive data may memorize sensitive personal information and photos, posing privacy and copyright concerns. |
| Approach: | They propose a framework that learns a universal noise pattern to recover unlearned information from MLLMs. |
| Outcome: | The proposed framework learns a universal noise pattern and can reveal unlearned content when applied to images. |
Copied to clipboard
| Challenge: | Existing chart understanding benchmarks focus on single-chart tasks, neglecting multi-hop reasoning required to extract and integrate information from multiple charts. |
| Approach: | They propose a benchmark that evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
| Outcome: | The proposed benchmark evaluates MLLMs’ capabilities in four key areas: direct question answering, parallel question answering and comparative reasoning. |
Copied to clipboard
| Challenge: | Existing benchmarks for video comment art are constrained by their limited modalities and insufficient categories, hindering creativity in video-based comment art creation. |
| Approach: | They propose a benchmark that integrates video and text modalities to evaluate MLLMs’ abilities to compose video Comment art. |
| Outcome: | The proposed framework integrates video and text modalities to evaluate MLLMs’ abilities to compose video comment art. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are hindered by the rapid growth of key–value (KV) caches. |
| Approach: | They propose a hybrid KV cache compression framework that reduces KV memory by 7.9 and speeds up decoding by 1.52. |
| Outcome: | Experiments on 11 multimodal benchmarks show that HYBRIDKV cuts KV cache memory by 7.9 and speeds up decoding by 1.52. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have seen growing adoption across various scientific domains. |
| Approach: | They propose a framework that bridges the molecule-text modality gap by integrating a comprehensive benchmark of pretraining strategies and dataset configurations. |
| Outcome: | The proposed framework improves multimodal LLMs through cross-modal alignment and multi-graph understanding. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models fine-tuned with multimodal instruction-following data have demonstrated formidable capabilities in multimodal tasks. |
| Approach: | They propose to employ four PEFT methods to fine-tune the LLM component of open-source MLLMs. |
| Outcome: | The proposed method is the best performing on seven datasets, while fine-tuning the connector layers leads to improved performance in most MLLMs. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data. |
| Approach: | They propose a method to disentangle risks through step-by-step reasoning within multimodal inputs. |
| Outcome: | The proposed approach improves safety alignment in MLLMs by fine-tuning and iterative Reinforcement Learning from AI feedback. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) lack understanding of multi-image and interleaved inputs due to the visual features encoded by frozen encoders before being fed into the LLM backbone. |
| Approach: | They propose a two phase paradigm to enable in-depth multimodal context fusion prior to feeding the features into LLMs. |
| Outcome: | The proposed paradigm boosts the performance on 7 multi-image scenarios, contributing to increments on average accuracy by 2.13% and 7.60% against strong MLLMs baselines with 3B and 11B LLMs, respectively. |
Copied to clipboard
| Challenge: | Existing methods for creating versatile MLLMs rely on joint training with paired instruction data, which is resource-intensive and challenging to extend to new modalities. |
| Approach: | They propose a new paradigm for multimodal large language models by reusing modality encoders and merging LLM parameters. |
| Outcome: | The proposed model retains the modal understanding capabilities of each original model. |
Copied to clipboard
| Challenge: | Document Image Machine Translation (DIMT) faces generalization challenges due to limited training data and the complex interplay between visual and textual information. |
| Approach: | They propose a single-to-mix Modality alignment framework leveraging Multimodal Large Language Models (MLLMs) this framework aligns an imageonly encoder with multimodal representations of an MLLM pre-trained on large-scale document image datasets. |
| Outcome: | The proposed framework improves translation quality in cross-domain generalization and challenging document image scenarios. |
Copied to clipboard
| Challenge: | Existing studies on English-centric translation tasks have focused on multimodal large language models, but the exploration of many-to-many translation is limited by the scarcity of parallel data. |
| Approach: | They propose a three-stage curriculum learning strategy that leverages the machine translation capabilities of large language models and adapts them to S2TT tasks. |
| Outcome: | The proposed strategy achieves state-of-the-art average performance in 1514 language pairs, requiring fewer than 10 hours of speech data per language to achieve competitive results. |
Copied to clipboard
| Challenge: | Existing benchmarks for MU are limited by a lack of image diversity and coarse-grained unlearning targets. |
| Approach: | They propose a benchmark to evaluate misinformation unlearning in MLLMs . OFFSIDE supports advanced unlearning targets such as fine-grained unlearning and visual rumor removal. |
| Outcome: | OFFSIDE supports advanced unlearning targets, such as fine-grained unlearning and visual rumor removal. |
Copied to clipboard
| Challenge: | Short-video platforms have become major channels for misinformation, but their robustness against misinformation entangled with cognitive biases remains under-explored. |
| Approach: | They propose a framework for evaluation of short-video platforms that use visual cues and social cue. |
| Outcome: | The proposed framework evaluates MLLMs across five modality settings. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models struggle with visual reasoning, despite strong performance on vision-language tasks. |
| Approach: | They propose a visually cued chain-of-thought prompting that enhances multi-step mathematical reasoning by explicitly referencing visual annotations in diagrams. |
| Outcome: | The proposed model improves GPT-4o's accuracy on an irregular polygon side-counting task from 7% to 93%. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models struggle with 3D spatial reasoning as they fail to construct structured abstractions of the 3D environment depicted in video inputs. |
| Approach: | They propose a prompting method that induces MLLMs to generate 3D representations as reasoning traces for more accurate spatial question answering. |
| Outcome: | Extensive experiments on VSI-Bench and OST-Bech show that TRACE improves over prior prompting strategies across a diverse range of MLLM backbones. |
Copied to clipboard
| Challenge: | Existing methods to identify multimodal neurons in MLLMs are insufficiently understood . previous studies focused on identifying neurons corresponding to single-tokens . |
| Approach: | They propose a method to identify multimodal neurons in Transformer-based MLLMs . they introduce fuzzy set theory to model the complex relationship between neurons and semantic concepts . |
| Outcome: | The proposed method improves performance on the Visual Question Answering task. |
Copied to clipboard
| Challenge: | Gradient Ascent (GA) has emerged as a promising approach for concept unlearning in Multimodal Generative Models (MGMs). |
| Approach: | They propose a novel approach that selectively applies GA to targeted Conceptual Knowledge while preserving Natural Knowledge through Gradient Descent (GD). |
| Outcome: | The proposed approach removes Conceptual Knowledge and inadvertently diminishes Natural Knowledge, resulting in utility degradation. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) excel in tasks ranging from image captioning to complex reasoning. |
| Approach: | They propose a contrastive decoding framework that dynamically calibrates each token generation by mining the model’s internal perceptual discrepancies. |
| Outcome: | The proposed framework mitigates hallucination while enhancing general reasoning capabilities. |
Copied to clipboard
| Challenge: | Visually Rich Document Understanding (VRDU) frameworks are a key area of research . early approaches to VRDU relied on manually crafted rules and domain-specific heuristics . conventional deep learning approaches do not integrate the diverse modalities in documents . |
| Approach: | They review recent advances in MLLM-based Visually Rich Document Understanding (VRDU) their findings highlight emerging trends and promising research directions . |
| Outcome: | The proposed frameworks are scalable, reliable, and adaptable, the authors argue . their findings highlight emerging trends and promising research directions . |
Copied to clipboard
| Challenge: | Existing metrics for conditional image generation are opaque and lack explainability . evaluators of these metrics have limited ability to evaluate image synthesis tasks . |
| Approach: | They propose a Visual Instruction-guided Explainable metric for evaluating conditional image models. |
| Outcome: | The proposed model achieves a high Spearman correlation with human evaluations, but is weaker than GPT-4o and GPT-v in evaluating synthetic images. |
Copied to clipboard
| Challenge: | Existing methods for visual token pruning compromise the integrity of visual understanding in pursuit of efficiency. |
| Approach: | They propose a model-agnostic method that integrates visual saliency and text relevance to reconcile efficiency with understanding by integrating visual salions and text relevant. |
| Outcome: | The proposed method outperforms state-of-the-art methods on LLaVA-NeXT . it achieves 13 decrease in FLOPs while maintaining 97% of original performance . |
Copied to clipboard
| Challenge: | Recent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities. |
| Approach: | They propose a simple MCL paradigm that addresses forgetting and misalignment . they propose 'MErge then ReAlign' to extend existing models to more modalities . |
| Outcome: | The proposed paradigm is easy to deploy and highly reusable in the MLLM community. |
Copied to clipboard
| Challenge: | MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. |
| Approach: | They propose to use Chinese images to assess MLLMs' higher-order perception and understanding of Chinese visual content. |
| Outcome: | The proposed model incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, to ensure the authenticity of the Chinese context. |
Copied to clipboard
| Challenge: | Multimodal representation is crucial for E-commerce tasks such as identical product retrieval. |
| Approach: | They propose an approach which leverages the generative power of Multimodal Large Language Models to extract key attributes from product images and text and enhances representation learning through a two-stage training framework. |
| Outcome: | The proposed model achieves state-of-the-art on multiple downstream retrieval tasks, validating the effectiveness of harnessing generative models to advance fine-grained representation learning. |
Copied to clipboard
| Challenge: | Existing methods for learning from errors lack a structured framework for analyzing and mitigating errors, especially in Multimodal Large Language Models (MLLMs). |
| Approach: | They propose a teacher-student framework that systematically structures errors to deliver targeted feedback for multimodal reasoning. |
| Outcome: | The proposed framework improves inference efficiency, token usage, and scalability by building a query-based structure that prioritizes visual information, diagnoses failure points, and guides corrective actions. |
Copied to clipboard
| Challenge: | Existing methods for MU forget quality and model utility are not fully explored for safety in MLLMs. |
| Approach: | They propose a safety unlearning benchmark for MLLMs to measure over-forgetting . they propose MU methods to forget quality and model utility . |
| Outcome: | The proposed method reduces over-forgetting by 79.5% while maintaining forget quality and model utility. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) enhance visual tasks by integrating visual representations into large language models. |
| Approach: | They propose a method to re-balance modalities by steering visual representations . they propose LLaVA Steering, a platform that enables rapid customization of MLLMs a component-based architecture . |
| Outcome: | The proposed model re-balances the modalities of visual representations in large language models . the model requires 500 times fewer trainable parameters than LoRA while maintaining comparable performance . |
Copied to clipboard
| Challenge: | Existing MCoT methods focus on inter-object reasoning, overlooking intra-object understanding crucial for image classification. |
| Approach: | They propose a Weak-supervision-guided Step-by-step Explanation method that reformulates MCoTs under weak supervision into concise, interpretable reasoning chains. |
| Outcome: | The proposed method improves interpretability by 37% and improves classification accuracy. |
Copied to clipboard
| Challenge: | generating accurate and faithful multimodal summaries is challenging due to lack of appropriate multimodal datasets . large language models excel at synthesizing key information from diverse sources, but lack of adequate multimodal data sets for fine-tuning . |
| Approach: | They propose a dataset specifically designed for image-text multimodal summarization . they generate summaries from Wikipedia sections and corresponding images and evaluate them . |
| Outcome: | The proposed dataset improves summary quality by training a critic model on human annotations and using its predictions to remove low-quality summaries. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are a promising tool for traditional education but lack authentic and domain-specific benchmarks to accurately interpret student handwritten solutions. |
| Approach: | They propose to use MLLMs to interpret unconstrained STEM student handwritten solutions with intertwined mathematical formulas, diagrams, and textual reasoning to bridge this gap. |
| Outcome: | The proposed model can detect and rectify recognition errors with minimal human intervention on unseen student solutions. |
Copied to clipboard
| Challenge: | Existing multimodal benchmarks overlook linguistic and visual ambiguities, authors say . ambiguity resolution between modalities is lacking in multimodal large language models . |
| Approach: | They propose a benchmark to evaluate multimodal ambiguity resolution across multilingual and cross-modal scenarios. |
| Outcome: | a new benchmark evaluates multimodal ambiguity resolution across multilingual and cross-modal scenarios . the benchmark shows that MLLMs can resolve ambiguities in image-text alignment . however, existing benchmarks often overlook linguistic and visual ambiguties . |
Copied to clipboard
| Challenge: | Existing benchmarks focus on evaluating MLLMs’ pre-existing knowledge or perceptual understanding, often neglecting the critical capability of reasoning. |
| Approach: | They propose a benchmark designed for visual clue-driven reasoning in daily scenarios that combines rigorous grounding in authentic daily activities and challenging query design that necessitates more than surface-level perception. |
| Outcome: | The proposed benchmark identifies visual clues and their ability to provide robust reasoning in daily scenarios. |
Copied to clipboard
| Challenge: | Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. |
| Approach: | They propose to use multiple-choice questions (MCQs) to assess LMs' understanding of traditional sports across 60 countries and 6 continents. |
| Outcome: | The new benchmark will be publicly available, fostering research in culturally aware AI systems. |
Copied to clipboard
| Challenge: | Mobile Agents are a key component of the “Agentic Economy” where they perform high-stakes financial transactions. |
| Approach: | They propose a systemic vulnerability termed Visual Dominance Hallucination (VDH) VDH exploits the modality gap in CLIP-based encoders via a novel Semantic-Decoupling Loss. |
| Outcome: | The proposed framework exploits the modality gap in CLIP-based encoders by preserving fidelity. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are susceptible to jailbreak attacks, authors say . multimodal information increases the risk of attacks, but also provides additional data . |
| Approach: | They propose a jailbreaking detector that detects maliciously perturbed image inputs . cross-modality information detector is designed to detect cross-modal similarity between harmful queries and adversarial images. |
| Outcome: | a new tool can detect maliciously perturbed image inputs without modification or computation cost. |
Copied to clipboard
| Challenge: | Recent advances in large language models have led to the development of multimodal large language model. |
| Approach: | They present a review of recent visual-based Large Language Models and analyze their architectures and alignment strategies. |
| Outcome: | The proposed models can integrate visual and textual modalities while providing a dialogue-based interface and instruction-following capabilities. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have expanded the capabilities of traditional language models by enabling interaction through both text and images. |
| Approach: | They propose a multimodal safety awareness benchmark to evaluate MLLMs across 29 safety scenarios with 1,500 carefully curated image-prompt pairs. |
| Outcome: | The proposed model is able to identify unsafe content and avoid over-sensitivity that can hinder helpfulness. |
Copied to clipboard
| Challenge: | Egocentric AI agents rely on pointing to resolve referential ambiguities in natural language commands. |
| Approach: | They propose a question-answering benchmark to evaluate and enhance pointing reasoning in egocentric views. |
| Outcome: | The proposed benchmark evaluates and enhances pointing reasoning in egocentric views. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are known to hallucinate, which limits their practical applications. |
| Approach: | They propose a method that uses three types of preference pairs to target hallucinations from their diverse forms and causes. |
| Outcome: | The proposed method surpasses most state-of-the-art methods and shows potential for further improvements. |
Copied to clipboard
| Challenge: | Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various vision-language tasks. |
| Approach: | They propose a systematic taxonomy to evaluate MLLMs' ability to interpret real-world music scores and answer complex musicological queries. |
| Outcome: | The proposed model is based on real-world music scores and user-generated questions and discussions, and is scalable and controlled. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are developing but lack external feedback . there is no clear on how to select reward models for agents . |
| Approach: | They propose a benchmark to evaluate agent reward modeling ability in MLLMs . they use multiple dimensions and real-world agent scenarios evaluation . |
| Outcome: | The proposed benchmark evaluates agent performance in multimodal large language models . it covers perception, planning, and safety with 7 scenarios and is highly difficult and high-quality . |
Copied to clipboard
| Challenge: | Existing approaches to mitigating vision-knowledge conflict in Large Language Models (MLLMs) are not effective and can be further scaled. |
| Approach: | They propose a framework to generate inputs to simulate and evaluate vision-knowledge conflict in Multimodal Large Language Models (MLLMs) using original images and 1,122 high-quality question-answer pairs, they propose 'a diagnostic benchmark' |
| Outcome: | The proposed framework, benchmark, and analysis contribute to the understanding and mitigation of vision-knowledge conflicts in Multimodal Large Language Models (MLLMs). |
Copied to clipboard
| Challenge: | Existing GUI reasoning methods rely on direct screen-based decision-making, which lacks interpretability and overlooks a comprehensive understanding of UI elements, ultimately leading to task failure. |
| Approach: | They propose a GUI reasoning paradigm that treats the GUI reasoning task as a cyclic ***Screen-UI elements-Action** process. |
| Outcome: | The proposed paradigm achieves state-of-the-art UI understanding performance while yielding superior results in GUI reasoning tasks. |
Copied to clipboard
| Challenge: | Existing static image-text benchmarks are insufficient for evaluating multimodal large language models’ dynamic perception and interactive reasoning abilities. |
| Approach: | They propose a game-based evaluation framework to assess multimodal large language models’ visual reasoning in dynamic, continuous-space environments. |
| Outcome: | The proposed framework systematically assesses MLLMs’ visual reasoning in dynamic, continuous-space environments. |
Copied to clipboard
| Challenge: | Existing efforts to mitigate this via token compression fail due to its autoregressive nature . linguistically redundant tokens are erroneously pruned, leading to hallucinations . |
| Approach: | They propose a method that reformulates token pruning as a Visual-Anchored Information Bottleneck (VA-IB) optimization problem. |
| Outcome: | Experiments on Qwen2-VL and Llama-3.2 families show that the proposed model achieves a speedup with negligible accuracy loss. |
Copied to clipboard
| Challenge: | Recent advances in Multimodal Large Language Models have raised serious safety concerns. |
| Approach: | They propose a method for manipulating the output preference of MLLMs using a preference hijacked image. |
| Outcome: | The proposed method works at inference time and requires no model modifications. |
Copied to clipboard
| Challenge: | Existing evaluation frameworks for multimodal large language models suffer from limitations . modality shortcuts and biased reasoning paths are common in such models . |
| Approach: | a new benchmark evaluates omni-modal multi-hop reasoning using 6,144 questions . authors propose OMHBench to address these limitations by comparing modalities . |
| Outcome: | OMHBench evaluates omni-modal multi-hop reasoning on 6,144 questions with balanced reasoning paths . evaluation of 13 state-of-the-art models shows performance gap exists between MLLMs and open-source models . |
Copied to clipboard
| Challenge: | Existing input-centric solutions fail to reverse this intrinsic mechanism of information loss. |
| Approach: | They propose a Variational Information Flow framework that leverages a probabilistic perspective to model visual saliency relevant to the question-answer pair as a latent distribution. |
| Outcome: | The proposed framework improves general VQA, fine-grained perception and visual grounding. |
Copied to clipboard
| Challenge: | Existing methods to improve difficulty calibration for Multimodal Large Language Models only consider text input . visual embeddings in training data reduce effectiveness of these methods . |
| Approach: | They propose a method to detect member samples in poorly generalized local manifolds by visual embeddings. |
| Outcome: | The proposed method surpasses existing methods. |
Copied to clipboard
| Challenge: | Current paradigms rely on holistic scoring and static leaderboards to disentangle fine-grained competencies. |
| Approach: | They propose a framework to shift the focus from ranking to fine-grained diagnosis. |
| Outcome: | The proposed framework surpasses the strongest baseline by 7.92%. |
Copied to clipboard
| Challenge: | Current physics benchmarks focus on text-only inputs or only on problem-solving . current physics reasoning benchmarks neglect critical intermediate steps of variable identification and process formulation. |
| Approach: | a new benchmark evaluates multimodal large language models in physics reasoning . the benchmark measures variables, process formulations, and solution derivation . |
| Outcome: | PhysicsArena is the first multimodal physics reasoning benchmark . it evaluates MLLMs across three critical dimensions: variable identification, process formulation, and solution derivation. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) show promising results in complex, human-centered environments, yet evaluating their capacity for nuanced, humanlike reasoning and decision-making remains challenging. |
| Approach: | They introduce VIVA+, a cognitively grounded benchmark for evaluating the reasoning and decision-making of MLLMs in human-centered situations. |
| Outcome: | The VIVA+ model is based on 1,317 real-world situations paired with 6,373 multiple-choice questions . it consists of three core abilities for decision-making: (1) Foundational Situation Comprehension, (2) Context-Driven Action Justification, and (3) Reflective Reasoning. |
Copied to clipboard
| Challenge: | Existing literature on visual storytelling has not explored the ideation process fully. |
| Approach: | They propose a visual story ideation task that automates the selection and arrangement of visual assets into coherent sequences that convey expressive storylines. |
| Outcome: | The proposed framework surpasses baseline by 33.5% and 18.5%, respectively, on three metrics. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal tokens. |
| Approach: | They propose a training-free pruning framework that prunes multimodal tokens without a trained pruning method. |
| Outcome: | The proposed pruning framework outperforms existing token pruning methods and generalizes across diverse MLLMs. |
Copied to clipboard
| Challenge: | Existing automated layout models are ill-suited for spreadsheets, authors say . existing layout models treat components as rectangles with continuous coordinates . authors: spreadsheets are powerful tools for organizing and analyzing data . |
| Approach: | They formalize a spreadsheet layout generation task and introduce a framework for spreadsheet layouts . they use multimodal large language models to combine rule and vision reflection . |
| Outcome: | The proposed framework outperforms baselines in a spreadsheet layout generation task by 22.6%. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have been a key advance in video understanding but their vulnerability to adversarial tampering remains underexplored. |
| Approach: | They evaluate MLLMs against five prevalent tampering techniques to assess their robustness . they use a tampered video format to examine the vulnerability of ML models . |
| Outcome: | The benchmark evaluates MLLMs against five prevalent tampering techniques based on 19 video manipulation tasks. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models (MLLMs) are predominantly trained on consistent visual-textual inputs, leaving open the question of whether they can handle semantic mismatches in layout-rich content. |
| Approach: | They propose to use multimodal inconsistency reasoning to assess MLLMs' ability to reason about semantic mismatches in webpages, presentation slides, and posters. |
| Outcome: | The proposed model outperforms open-source models in detecting inconsistencies in webpages, presentation slides, and posters while remaining vulnerable to inconsistent errors. |
Copied to clipboard
| Challenge: | Existing safety benchmarks fail to provide reliable assessments due to limited risk coverage, insufficient scale and the oversight of complex modality combinations. |
| Approach: | They propose a framework that covers 61 risk categories across four modality interactions to address this gap. |
| Outcome: | The proposed framework covers 61 risk categories across four distinct modality interactions. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have demonstrated remarkable performance across various tasks, effectively following instructions to meet diverse user needs. |
| Approach: | They propose a framework for evaluation benchmarks and attack techniques for LLMs and MLLMs to enhance their security. |
| Outcome: | The proposed frameworks have been exploited to exploit the weaknesses of LLMs and MLLMs. |
Copied to clipboard
| Challenge: | Existing methods for scene graph generation lack task-specific structured reasoning and sparse, long-tailed relation distributions. |
| Approach: | They propose a structured reasoning framework that integrates task-specific Chain-of-Thought and reinforcement learning with group sequence policy optimization to achieve unbiased scene graph generation. |
| Outcome: | The proposed framework achieves superior performance on two benchmarks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models excel in general tasks but struggle with specialized, structured cultural symbols. |
| Approach: | They evaluate 21 leading MLLMs and compare their performance to a benchmark for Ancient Chinese musical notation. |
| Outcome: | The benchmark evaluates 21 leading MLLMs on five types of ancient Chinese music notation systems. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are increasingly being deployed as content moderators . however, they exploit the Human-AI capability gap and create adversarial environments . smuggling attacks exploit the human-AI gap and exploit the vulnerability . |
| Approach: | They construct a benchmark to evaluate the vulnerability of MLLMs as content moderators . they identify three root causes: limited capabilities of vision encoders, robustness gap in OCR . |
| Outcome: | The proposed model exploits the Human-AI capability gap and is vulnerable to smuggling attacks. |
Copied to clipboard
| Challenge: | State-of-the-art multimodal web agents can perform many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs). |
| Approach: | They propose to build multimodal web agents for few-shot adaptability using human demonstrations to improve their generalization and adaptability. |
| Outcome: | The proposed framework enables both proprietary and open-weights multimodal web agents to adapt to new websites and domains using few human demonstrations. |
Copied to clipboard
| Challenge: | a new framework for image-text instruction data evolution improves MLLM performance . lack of high-quality instruction data remains a major bottleneck in ML modeling . |
| Approach: | They propose a multimodal instruction data evolution framework that iteratively enhances data quality through fine-grained perception, cognitive reasoning, and interaction evolution. |
| Outcome: | The proposed approach improves MLLM performance in nine vision-language tasks while using significantly less data. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have shown promise in MER, but their internal decision-making mechanisms under modality conflict and missingness remain underexplored. |
| Approach: | They propose a multimodal large language model that can detect and control modality conflicts and missing subsets by a lightweight mechanism that detects and controls modality conflict. |
| Outcome: | The proposed framework improves performance across settings, showing it can handle conflict and missing behaviors. |
Copied to clipboard
| Challenge: | Effective reward modeling is especially valuable in reinforcement learning (RLHF) . |
| Approach: | They propose a paradigm for empowering general-purpose MLLMs judges with strong reasoning capabilities by using multiple-choice problem models instead of directly assigning scores. |
| Outcome: | The proposed model surpasses GPT-4o on VL-RewardBench and improves performance on MM-Vet by up to 7.7%. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on assessing factual and logical correctness in downstream tasks with limited emphasis on evaluating MLLMs’ ability to interpret pragmatic cues and intermodal relationships. |
| Approach: | They propose to use Coherence Relations to assess MLLMs' ability to perform multimodal discourse analysis using different prompting strategies. |
| Outcome: | The proposed model fails to match the performance of simple classifier-based benchmarks on 10+ MLLMs using different prompting strategies. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, but precise coordinate prediction remains a challenge. |
| Approach: | They propose a training-free, inference-time correction method to correct VPEs . they isolate position-unconditioned tendencies by shuffling VPE and use it to steer digit decoding . |
| Outcome: | The proposed method is training-free, inference-time correction method . it effectively rectifies coordinate drift, yielding consistent improvements without retraining . |
Copied to clipboard
| Challenge: | Existing studies have explored the evolutionary analysis of ancient scripts, with particular attention to the transformation of character forms from oracle bone inscriptions to regular script. |
| Approach: | They propose a benchmark framework that leverages MLLMs to analyze the evolution of ancient Chinese scripts. |
| Outcome: | The proposed framework improves performance on core tasks and character recognition and evolutionary reasoning tasks while limiting performance on other tasks. |
Copied to clipboard
| Challenge: | Multimodal instruction fine-tuning degrades textual reasoning capability, undermining multimodal performance. |
| Approach: | They propose a plateau-guided model merging method that selectively injects base language model parameters into MLLMs to mitigate this degradation. |
| Outcome: | The proposed framework reduces multimodal instruction fine-tuning degradation by incorporating a plateau-guided model merging method into MLLMs. |
Copied to clipboard
| Challenge: | Advanced GUI agents suffer from prohibitive deployment costs on resource-constrained devices. |
| Approach: | They propose a lightweight GUI agent with GUI-specific knowledge and task scalability . LAMO-3B supports monolithic execution and MAS-style orchestration . |
| Outcome: | The proposed GUI agent LAMO-3B supports monolithic execution and MAS-style orchestration. |
Copied to clipboard
| Challenge: | Empirical evaluations on state-of-the-art MLLMs reveal a significant performance gap . ML models lack the fine-grained cross-modal reasoning required to bridge visual discontinuities. |
| Approach: | They propose a benchmark that renders fragmented documents directly from Markdown to facilitate evaluation of VRDU tasks. |
| Outcome: | The proposed benchmark renders fragmented documents directly from Markdown. |
Copied to clipboard
| Challenge: | Existing benchmarks for Geometry problem solving lack fine-grained evaluation for long-step problems necessitating auxiliary line construction. |
| Approach: | They present a fine-grained annotated dataset with long-step reasoning and auxiliary line construction that provides a detailed evaluation of 23 leading MLLMs. |
| Outcome: | The proposed model performs significantly worse on long-step problems than short-step ones, with 18 models showing a performance drop of over 50%. |
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of MLLMs to generate single tokens one by one, while lacking studies about how their representation vectors can encode global multimodal information. |
| Approach: | They propose to use image-caption corpus to train Multimodal Large Language Models (MLLMs) . they find that the topmost layers encode more global semantic information . |
| Outcome: | The proposed models can encode more global semantic information, rather than the topmost layers, and perform better on visual-language entailment tasks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are increasingly used as automatic judges . however, their reliability and vulnerabilities to biases remain underexplored . |
| Approach: | They propose a benchmark to evaluate MLLMs that fail to integrate visual cues . they also introduce a test to evaluate the reliability of MLMLs based on a set of asymmetric evaluation tendencies. |
| Outcome: | Experiments on 26 state-of-the-art MLLMs reveal modality neglect and asymmetric evaluation tendencies . a standardized model with a benchmark enables a fine-grained diagnosis of nine bias types . |
Copied to clipboard
| Challenge: | Existing benchmarks focus on simple image-text interactions, overlooking complex visual formats like charts. |
| Approach: | They propose a semi-automatic framework for generating evaluation samples through multi-modal keypoint extraction, knowledge graph construction, and qa pair synthesis. |
| Outcome: | The proposed framework generates 4,738 question-answering pairs across 8 domains from real-world documents. |
Copied to clipboard
| Challenge: | Existing benchmarks for document understanding in the wild are based on scanned or digital documents . however, these benchmarks fail to capture the challenges posed by documents in the real world . |
| Approach: | They propose a new benchmark that incorporates a diverse set of manually captured document images reflecting real-world conditions. |
| Outcome: | The proposed model is based on a set of manually captured document images reflecting real-world conditions and is compared with digital or scanned documents. |
Copied to clipboard
| Challenge: | Existing methods for MLLMs are weak on explicit attacks, but weak on implicit ones. |
| Approach: | They propose an automated red-teaming pipeline that leverages reinforcement learning with tailored reward modules to generate diverse implicit samples across 14 domains. |
| Outcome: | The proposed method outperforms existing methods in implicit and explicit attacks while maintaining high utility. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) outperform existing benchmarks in both natural language and coding domains. |
| Approach: | They propose a scalable benchmark that integrates vision and language modalities to address this gap by eliminating textual shortcuts. |
| Outcome: | The new benchmark outperforms existing benchmarks in both natural language and coding domains. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models struggle with dynamic interactions due to the scarcity of high-quality interleaved data. |
| Approach: | They propose a large-scale interleaved live interaction Chinese dataset with human-annotated video responses. |
| Outcome: | The proposed model can be used to evaluate live interactions in Chinese over 1,100 hours and 80,037 dialogue turns. |
Copied to clipboard
| Challenge: | Current methods struggle to distinguish targets in low Signal-to-Noise Ratio environments and lack sufficient pre-execution verification to prevent error accumulation. |
| Approach: | They propose a Memory-augmented Debate System to ensure precise grounding across diverse interfaces and handle irreversible errors in extended workflows. |
| Outcome: | The proposed system achieves a 90.23% task success rate on MaDS-Benchmark and strong performance on public benchmarks including AITW, AITZ, CAGUI, and GUIOdyssey. |
Copied to clipboard
| Challenge: | Existing low-resource security alignment methods struggle with the security risks posed by additional modalities. |
| Approach: | They propose to use multimodal datasets to enhance safety alignment but it is costly to construct these datasets. |
| Outcome: | Experiments on image, video, and audio-based MLLMs show that the proposed method can synthesize a high-quality embedding on a single RTX3090 GPU within 24 seconds. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have shown strong performance in document image tasks, especially Optical Character Recognition (OCR). However, they struggle with Document Image Machine Translation (DIMT), which requires handling both cross-modal and cross-lingual challenges. |
| Approach: | They propose a novel fine-tuning paradigm that allows the model to generate OCR text before producing translation text, which allows it to leverage its strong monolingual OCR ability while learning to translate text across languages. |
| Outcome: | The proposed model can leverage its strong monolingual OCR ability while learning to translate text across languages. |
Copied to clipboard
| Challenge: | Current mathematical benchmarks focus on evaluating MLLMs’ problem-solving ability, yet there is a crucial gap in addressing more complex scenarios such as error detection. |
| Approach: | They propose to evaluate multimodal error detection by evaluating two sub-tasks error step identification and error categorization. |
| Outcome: | The proposed task evaluates MLLMs' ability to handle multimodal questions compared to text-only models. |
Copied to clipboard
| Challenge: | Existing approaches to training GUI agents on dynamic tasks are based on SFT or Behavior Cloning. |
| Approach: | They propose a framework that integrates global trajectory insights directly into offline learning . they reconstruct diverse rollout candidates from static data and detect first failure point . |
| Outcome: | The proposed framework improves long-horizon task completion rates and robustness compared to baselines. |
Copied to clipboard
| Challenge: | Current scientific reasoning models struggle with generalization across domains and fall short of multimodal perception. |
| Approach: | They propose to use multimodal large language models to integrate text, images, and other modalities to enhance scientific reasoning. |
| Outcome: | The proposed models can integrate text, images, and other modalities and improve reasoning across disciplines. |
Copied to clipboard
| Challenge: | a new data-centric approach could address cultural gaps in multimodal large language models . despite being trained on billions of image-text pairs, today's models are biased towards English and Western data. |
| Approach: | They propose a data-centric approach that directly grounds MLLMs in cultural knowledge. |
| Outcome: | The proposed approach outperforms open-source models on cultural-focused benchmarks without degrading results on mainstream vision–language tasks. |
Copied to clipboard
| Challenge: | Prior studies have focused on strengthening multimodal reasoning by improving representation alignment or increasing computation, but these methods do not characterize the differences in visual demands across tasks. |
| Approach: | They propose an entropy-driven task-adaptive visual attention allocation framework that uses visual attention entropic as a control signal to dynamically allocate attention according to task demands. |
| Outcome: | The proposed framework achieves consistent performance gains across diverse reasoning tasks, datasets, and models, providing a clear direction toward more reliable multimodal reasoning. |
Copied to clipboard
| Challenge: | Existing benchmarks focus on single image settings, but some focus on multi-image settings. |
| Approach: | They introduce the TempVS benchmark which focuses on temporal grounding and reasoning capabilities of Multimodal Large Language Models in image sequences. |
| Outcome: | The proposed model performs poorly compared to human models in vision and language tasks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models perform well on visual question answering tasks, but it remains unclear whether their reasoning relies more on memorized world knowledge or on visual information present in the input image. |
| Approach: | They propose a dataset of visual-realistic counterfactuals that put world knowledge priors into conflict with visual input. |
| Outcome: | The proposed dataset puts world knowledge priors into conflict with visual input . it shows that model predictions shift toward visual evidence in mid-to-late layers . |
Copied to clipboard
| Challenge: | Existing benchmarks for multimodal large language models are limited to multiview diagnostics. |
| Approach: | They propose a benchmark specifically designed for medical multi-image understanding that evaluates MLLMs across four dimensions. |
| Outcome: | The proposed model performs better in multi-image contexts than open-source models . the model perform better when processing increased visual loads than closed-source ones . |
Copied to clipboard
| Challenge: | Xu and Stone et al., 2014, show eye movements are correlated with discourse goals but the relationship between eye movements and coherence is a missing link. |
| Approach: | They propose an eye gaze pattern ranking algorithm and a semantic gaze visualization technique to study eye gaze patterns and coherence relations in multimodal language contexts. |
| Outcome: | The proposed method combines eye-tracking and a semantic gaze visualization technique to study eye movements in multimodal language contexts. |
Copied to clipboard
| Challenge: | Existing methods for short video fake news detection rely on black-box MSLMs with poor explainability and superficial understanding or on specific prompt strategies for Multimodal Large Language Models (MLLMs) |
| Approach: | They propose a multi-agent framework called CSI for short video fake news detection. |
| Outcome: | The proposed framework provides rigorous explanations while achieving state-of-the-art performance on two real-world datasets. |
Copied to clipboard
| Challenge: | Recent studies have explored Continual Instruction Tuning (CIT) in Multimodal Large Language Models (MLLMs), with a primary focus on Task-incremental CIT, where MLLM are required to continuously acquire new tasks. |
| Approach: | They propose a Sparse Mixture of Expert (SMoE) based method for domain-incremental CIT in Multimodal Large Language Models (MLLMs) . they equip the SMoA module with a domain-specific autoregressive loss (DSAL) they establish a new benchmark to evaluate the efficacy of their method . |
| Outcome: | The proposed method outperforms all baselines and is based on a Sparse Mixture of Experts (SMoE) module . |
Copied to clipboard
| Challenge: | Existing studies overlook the need of mining relations among multiple columns rather than just the semantic relation between two specific columns in real-world practice. |
| Approach: | They propose a Chain-of-Thought distillation framework with self-correction mechanism to enhance MLLMs’ reasoning capabilities without increasing parameter scale. |
| Outcome: | The proposed method significantly outperforms baselines on wide datasets. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models have demonstrated remarkable capabilities across vision-language tasks, but their performance as embodied agents needs further exploration. |
| Approach: | They propose a framework to evaluate multimodal large language models as zero-shot agents . they find that enhancing prevalent agents with Chain-of-Thought reasoning and self-reflection leads to an unexpected performance decrease. |
| Outcome: | The proposed framework enables comparisons and component-level ablations across diverse MLLM architectures, agent designs, and navigation tasks. |
Copied to clipboard
| Challenge: | MLLMs assume linguistic context invariably enhances visual understanding . a diagnostic benchmark is used to evaluate ML models under hierarchical linguistic interference . |
| Approach: | They propose a diagnostic benchmark to evaluate MLLMs under hierarchical linguistic interference. |
| Outcome: | The proposed benchmark compared 402 videos with a physical constraint set to evaluate MLLMs under hierarchical linguistic interference. |
Copied to clipboard
| Challenge: | Visual Instruction Tuning (VIT) aims to enhance Multimodal Large Language Models (MLLMs), but its effectiveness is often compromised by corrupted datasets with issues such as hallucinated content and poor OCR quality. |
| Approach: | They propose a corruption-robust training paradigm that surpasses existing strategies for mitigating the effects of corrupted data. |
| Outcome: | The proposed training paradigm surpasses existing strategies for mitigating the effects of corrupted data. |
Copied to clipboard
| Challenge: | Existing text-rich image understanding benchmarks lack scale and fragmented scenarios . a new full-image structured output format is proposed to enable fine-grained evaluation of perception and reasoning capabilities. |
| Approach: | They propose a large-scale, multilingual benchmark that includes over 100,000 annotations and 22,000 question-answer pairs. |
| Outcome: | The proposed framework provides a comprehensive platform for developing and evaluating next-generation multimodal AI systems. |
Copied to clipboard
| Challenge: | Existing multimodal large language models struggle when faced with unseen domains or languages. |
| Approach: | They propose a framework that leverages the broad knowledge of an MLLM to generate cross-modal pre-questions (preQs) before retrieval. |
| Outcome: | Experiments show that PREMIR outperforms existing retrievers on out-of-distribution benchmarks, including closed-domain and multilingual settings, outperforming strong baselines across all metrics. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Multimodal Large Language Modells (MLLMs) are increasingly being deployed across a range of domains, including finance, law, peer review, and recruitment. |
| Approach: | They investigated how image-based evaluations are influenced by non-job-related information, including extracurricular activities and social media images. |
| Outcome: | The proposed models exhibit significant halo effects in image-based evaluations while text-based assessments showed more resistance to bias. |
Copied to clipboard
| Challenge: | Existing multimodal large language models lack domain-specific expertise to perform chemical tasks. |
| Approach: | They propose a benchmark dataset for evaluating multi-step multimodal reasoning capacities in the chemistry domain. |
| Outcome: | The proposed model surpasses existing models in all CheMM-Bench tasks. |
Copied to clipboard
| Challenge: | Existing approaches to interleaved reasoning are limited by the cost of re-encoding pixel-dense images. |
| Approach: | They propose a framework that unifies dynamic state evolution with precise perceptual modeling. |
| Outcome: | The proposed framework outperforms existing approaches on multimodal reasoning benchmarks. |
Copied to clipboard
| Challenge: | Recent advances have enabled MLLMs to tackle complex challenges such as mathematical reasoning and multimodal understanding. |
| Approach: | They propose a multimodal refinement benchmark to evaluate the refinement capabilities of Multimodal Large Language Models (MLLMs) the benchmark categorizes errors into six error types to highlight areas for improvement in effective reasoning enhancement. |
| Outcome: | The proposed framework evaluates the refinement capabilities of multimodal large language models across six scenarios. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have experienced rapid development in recent years, but there is a notable lack of effective and specialized multimodal evaluation datasets in the financial domain. |
| Approach: | They introduce FinMME, a multimodal large language model with 11,000 financial research samples and 20 annotators. |
| Outcome: | The proposed model performs better than state-of-the-art models, highlighting its challenging nature. |
Copied to clipboard
| Challenge: | Mobile Phone Agents (MPAs) have attracted huge attention due to their practicability in a multitude of scenarios. |
| Approach: | They propose a data mixture optimization solution that extrapolates optimal data mixtures from a trainable network. |
| Outcome: | The proposed model outperforms existing methods on open-source benchmarks and on open source benchmarks. |
Copied to clipboard
| Challenge: | Existing prompting methods for multimodal large language models lack fine-grained perception across disparate images . existing methods fail to integrate perception and reasoning, causing problems with general multi-image reasoning tasks. |
| Approach: | They propose a generalized prompting method that integrates perception and reasoning . they evaluate the method on open-source and closed-source MLLMs . |
| Outcome: | The proposed method shows competitive performance across tasks and improves in challenging scenarios. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models excel at visual perception and reasoning in third-person and egocentric videos, but are prone to hallucinations, generating coherent yet inaccurate responses. |
| Approach: | They propose to use a benchmark to evaluate MLLM hallucinations in egocentric videos. |
| Outcome: | EGOILLUSION comprises 1,400 videos paired with 8,000 human-annotated open and closed-ended questions designed to trigger hallucinations in both visual and auditory cues in egocentric videos. |
Copied to clipboard
| Challenge: | MLLMs have facilitated multimodal summarization with multimodal outputs, but their evaluation is fragmented . MM-Eval integrates assessments of textual quality, cross-modal alignment, and visual diversity . |
| Approach: | They propose a unified evaluation framework that integrates assessments of textual quality, cross-modal alignment, and visual diversity. |
| Outcome: | The proposed framework improves over heuristic aggregation baselines and provides an interpretable, reference-weak framework for comparative evaluation of multimodal summaries. |
Copied to clipboard
| Challenge: | Existing approaches to debiase MLLMs rely on handcrafted prompts that are brittle and difficult to generalize across tasks and bias types. |
| Approach: | They propose an adaptive self-debiasing framework that optimizes task-specific debiasers to suppress stereotypical outputs. |
| Outcome: | The proposed framework suppresses stereotypical outputs while maintaining performance. |
Copied to clipboard
| Challenge: | Existing knowledge editing methods for MLLMs lack multi-granularity knowledge . existing knowledge editing approaches lack multimodality knowledge and generalize to multimodal data. |
| Approach: | They propose a multimodal knowledge editing method which integrates key knowledge layers within MLLMs and collaboratively edits them. |
| Outcome: | The proposed method improves visual generality performance on knowledge data of different granularities. |
Copied to clipboard
| Challenge: | Recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across various vision reasoning tasks. |
| Approach: | They propose a unified formal language that integrates plane and solid geometry, comprehensively covering geometric structures and semantic relations. |
| Outcome: | The proposed language achieves state-of-the-art parsing performance and significantly boosts MLLMs’ capabilities for downstream geometry reasoning tasks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) are powerful at integrating diverse data but struggle with complex reasoning. |
| Approach: | They propose a method which separates responses into positive and negative groups to stabilize training and preserve knowledge. |
| Outcome: | The proposed model View-R1 achieves a 10.55% improvement in reasoning and outperforms larger models while maintaining and improving performance on general tasks. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models (MLLMs) have advanced Chinese Classical Studies (CCS) but the audio dimension of CCS remains underexplored due to a lack of high-quality, domain-specific audio corpora. |
| Approach: | They propose a 119-hour audio corpus comprising 22,000 audio samples to bridge this gap . it encompasses a diverse range of literary genres across six tasks . |
| Outcome: | The proposed corpus encompasses a diverse range of literary genres across six tasks: Automatic Speech Recognition (ASR), Speech-to-Text Translation (S2TT), Speech Emotion Captioning (SEC), Spoken Question Answering ( SQA), Speech Understanding (SU), and Speech Reasoning (SR). |
Copied to clipboard
| Challenge: | Recent advances in Multimodal Large Language Models have significantly improved reasoning and generation tasks by leveraging joint vision-language representations. |
| Approach: | They propose a framework that reconciles inconsistencies across knowledge sources . they use a four-stage pipeline to generate an internal response from parametric knowledge . |
| Outcome: | Experiments on KB-VQA show that CoRe-MMRAG achieves performance gains of 5.6% and 9.3% over baseline methods. |
Copied to clipboard
| Challenge: | Recent work of GUI action grounding fine-tunes data from pre-trained MLLMs, but data is limited to specific GUI environments. |
| Approach: | They propose to use a GUI-based agent to collect environment-specific data and fine-tune GUI grounding models with the collected data. |
| Outcome: | The proposed model can be extended to other GUI environments to improve performance. |
Copied to clipboard
| Challenge: | Existing methods for cross-modal alignment assume a symmetric interaction between visual and textual modalities, implying that both spaces adapt to each other. |
| Approach: | They propose a method that regularizes the projector to maintain the geometric structure of the text embedding space via spectral filtering. |
| Outcome: | The proposed method preserves the LLM’s inherent linguistic capabilities and reduces object hallucination significantly better than standard fine-tuning methods. |
Copied to clipboard
| Challenge: | Recent advances in multimodal reasoning may pose new safety risks . evaluators neglect reasoningbased safety, where harm emerges only through MLLMs . |
| Approach: | They introduce a benchmark for multi-image reasoning safety that includes 2,676 instances . they find that models with more advanced multi- image reasoning are more vulnerable . |
| Outcome: | The proposed benchmark consists of 2,676 instances covering 9 multi-image relations . the results show that models with more advanced multi- image reasoning are more vulnerable . |
Copied to clipboard
| Challenge: | Current evaluations for Vision-language Models remain heavily anchored to ImageNet . |
| Approach: | They propose a large-scale semantically-annotated multimodal resource that extends the range of visual concepts, including diverse abstract categories. |
| Outcome: | The proposed model expands the range of visual concepts, including diverse abstract categories. |
Copied to clipboard
| Challenge: | This paper explores using Multimodal Large Language Models (MLLMs) to respond to student questions from online lectures . MLLM is a novel question answering task of real world significance . |
| Approach: | They propose to use Multimodal Large Language Models to automatically respond to student questions from online lectures by using a dataset of 5252 question-answer pairs from 296 computer science videos. |
| Outcome: | The proposed model can fine tune and fine tune questions from 296 computer science videos and show that students' preferences are important to the task. |
Copied to clipboard
| Challenge: | Existing approaches to overcome object hallucination are limited . Existing mitigations include costly retraining and a training-free inference framework . |
| Approach: | They propose a training-free inference framework that simulates a metacognitive self-correction process. |
| Outcome: | The proposed framework reduces object hallucination rates by 12.67% on MMHal-Bench and improves accuracy by 5.8% on POPE. |
Copied to clipboard
| Challenge: | Existing Chain-of-Thought (CoT) approaches lack intrinsic correction mechanisms, rendering them vulnerable to error propagation. |
| Approach: | They propose a multi-agent framework that enforces diagnostic rigor through adversarial dialectics. |
| Outcome: | Empirical evaluations show that the proposed framework improves explanation faithfulness and mitigates hallucinations. |
Copied to clipboard
| Challenge: | Existing debiasing methods create biased responses by completely removing an entire modality, forming an extreme and static training environment. |
| Approach: | They propose a method to debiase multimodal large language models by masking one modality and then enlarge the margin between clean and adversarial responses. |
| Outcome: | The proposed method achieves superior debiasing performance while maintaining general capabilities. |
Copied to clipboard
| Challenge: | Existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, but lack a critical gap in evaluating an agent. |
| Approach: | They evaluate multimodal large language models with six categories of kitchen hazards . they propose a safety-based approach that prioritizes multi-step corrective actions . |
| Outcome: | The proposed model can recognize hazards in QA settings, but average mitigation success rates are low . the proposed model is based on the embodied agent benchmark ALFRED . |
Copied to clipboard
| Challenge: | Existing benchmarks for Multimodal Large Language Models (MLLMs) have been lacking due to the rich nature of social interaction. |
| Approach: | They propose a video benchmark to evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction. |
| Outcome: | The proposed benchmarks evaluate MLLMs' capabilities across social scene understanding, social state reasoning, and social dynamics prediction tasks. |
Copied to clipboard
| Challenge: | Existing Multimodal Large Language Models (MLLMs) fail under RDEI, leading to disrupted structure and evidence-unsupported hallucinations. |
| Approach: | They propose a backbone-agnostic, evidence-driven pipeline that treats off-the-shelf MLLMs as interchangeable components to improve stem consistency and figure consistency. |
| Outcome: | The proposed pipeline improves stem consistency by 1.01-3.18%, figure consistency by 0.50-49.16%, and refusal F1 by 1.06-10.88% across question types. |
Copied to clipboard
| Challenge: | Existing evaluation metrics suggest that Multimodal large language models have acquired fine-grained visual grounding capabilities. |
| Approach: | They propose a benchmark to assess Referring Expression Comprehension (REC) that uses intra-image visual cues to localize target objects and a controllable evaluation mechanism to test sensitivity to fine-grained factual changes. |
| Outcome: | The proposed benchmarks show that multimodal large language models have a high level of performance on the RefCOCO family of benchmarks. |
Copied to clipboard
| Challenge: | Early approaches focus on text-based reasoning, but they often follow a single task-specific reasoning pattern. |
| Approach: | They propose a generative multimodal reasoning paradigm that unifies diverse reasoning skills by generating intermediate images during the reasoning process. |
| Outcome: | The proposed model unifies diverse multimodal reasoning skills by generating intermediate images during the reasoning process. |
Copied to clipboard
| Challenge: | Multimodal tables are ubiquitous in real applications but are difficult to evaluate in multimodal large language models. |
| Approach: | They propose a multimodal table benchmark that compares 500 real-world tables with 4021 question–answer pairs. |
| Outcome: | MMtabReal spans four question types, five reasoning categories, and eight structural archetypes. |
Copied to clipboard
| Challenge: | Recent studies focus on surface-level features, overlooking how design choices influence user behavior at scale. |
| Approach: | They propose a benchmark for multimodal understanding of how UI/UX design affects user behavior built on 300 real-world UI image pairs from industry A/B tests. |
| Outcome: | The proposed benchmarks show that models exhibit limited understanding of the behavioral impact of UI/UX design. |
Copied to clipboard
| Challenge: | Multimodal Large Language Models are increasingly deployed as social agents . yet their ability to integrate conflicting identity cues remains underexplored . |
| Approach: | They audit gender bias in MLLMs that pair synthetic voices with avatars of varying gender presentation and visual fidelity. |
| Outcome: | The findings show that multimodal fairness is not monolithic . they show that models may appear unbiased on one dimension while enforcing stereotypes on another . |
Copied to clipboard
| Challenge: | Multimodal Process Reward Models (MPRMs) have emerged as a pivotal framework for enhancing the reasoning capabilities of Multimodal Large Language Models. |
| Approach: | They propose a benchmark specifically designed to evaluate MPRMs’ proficiency in detecting erroneous reasoning steps across diverse error categories. |
| Outcome: | The proposed model achieves up to 4.8% performance improvement through test-time scaling. |
Copied to clipboard
| Challenge: | Existing models for visual information extraction suffer from limitations in scale and realism . ReceiptBench is a large-scale, human-annotated benchmark for receipts . |
| Approach: | They propose a large-scale, human-annotated benchmark for visual information extraction . the method organizes information extraction into four hierarchical sub-tasks . |
| Outcome: | The proposed method surpasses proprietary models on complex reasoning tasks. |
Copied to clipboard
| Challenge: | Existing approaches to GMNER use MLLMs as auxiliary tools, causing cumulative error propagation and a lack of rigorous cross-modal verification. |
| Approach: | They propose a model that enforces structured cross-modal reasoning through Multi-style Reasoning Schema Injection and Constraint-guided Verifiable Optimization. |
| Outcome: | The proposed model enforces structured cross-modal reasoning through multi-style Reasoning Schema Injection and Constraint-guided Verifiable Optimization. |
Copied to clipboard
| Challenge: | Existing unlearning methods suffer from a geometric mismatch, causing catastrophic forgetting or unsafe substitution. |
| Approach: | They propose a framework for surgical semantic pruning within the Lorentz manifold. |
| Outcome: | Experiments on MLLMU-Bench show that LOTUS significantly outperforms baselines while maintaining general utility. |
Copied to clipboard
| Challenge: | Existing methods for text-to-image alignment evaluation rely on coarse-grained metrics or static Question Answering pipelines that lack fine-grounded interpretability and struggle to reflect human preferences. |
| Approach: | They propose a reinforcement-guided visual reasoning framework for element-level text-to-image alignment evaluation. |
| Outcome: | The proposed framework achieves state-of-the-art results on four benchmarks and surpasses the strong proprietary Gemini 3 Pro and Training-based baselines. |